British Journal of Ophthalmology
● BMJ
Preprints posted in the last 90 days, ranked by how well they match British Journal of Ophthalmology's content profile, based on 14 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Samico, G. A.; Solages, N.; Scherer, R.; Muralidhar, R.; Gutkind, N. E.; Palazoni, V.; Medeiros, F. A.; Swaminathan, S. S.
Show abstract
Purpose: To evaluate the performance of secure cloud-based large language models (LLMs) in extracting glaucoma diagnosis, type, and severity from free-text clinical notes in the electronic health record (EHR). Design: Retrospective chart review analysis. Participants: 1,250 subjects from the Bascom Palmer Ophthalmic Repository. Methods: Clinical notes of glaucoma-related encounters between 2014 and 2024 were extracted from the Bascom Palmer Ophthalmic Repository. Two fellowship-trained glaucoma specialists annotated clinical notes for glaucoma presence, type, and severity at the eye level. The dataset was split into development (10%), validation (10%), and test (80%) sets. Development and validation sets were used for prompt engineering and refinement, and the held-out test set was used for evaluation. Five LLMs (Claude Opus 4.6, DeepSeek-V3.2, GPT-5.2, Grok 4.1, and Qwen3.6-35B-A3B) were accessed via Azure AI Foundry within HIPAA-compliant containers. Model performance was assessed using standard metrics. Clinician-entered ICD-10 codes were also compared with adjudicated labels. Main Outcome Measures: Gwet AC1, accuracy, sensitivity, specificity, and F1-score. Results: Inter-grader agreement was high for glaucoma detection (Gwet AC1= 0.930 (95% CI: 0.917-0.945), type classification (Gwet AC1= 0.917 (95% CI: 0.904-0.930), and severity staging (Gwet AC1= 0.901 (95% CI: 0.884-0.916). For glaucoma diagnosis, LLMs demonstrated high overall accuracy, with Claude achieving 97.5%, DeepSeek 96.0%, GPT 96.2%, Grok 94.4%, and Qwen 95.5%. F1 scores for glaucoma detection ranged from 95.4% to 98.9% across models. For glaucoma type classification, accuracies were 97.1%, 94.2%, 94.2%, 94.0%, and 94.4% for Claude, DeepSeek, GPT, Grok, and Qwen, respectively. F1 scores for the most prevalent type (POAG) ranged from 96.3% to 98.9%. For severity staging, accuracies were 95.0%, 94.8%, 94.5%, 94.0%, and 95.2%, respectively, with F1 scores ranging from 89.7% to 96.3% across severity categories and models. ICD-10 codes demonstrated substantially lower performance for type and severity staging, with overall accuracies of 89.2% and 58.5%, respectively. Conclusions: Secure cloud-based LLMs accurately extracted glaucoma diagnosis, type, and severity information from free-text ophthalmology notes, achieving performance approaching expert clinician adjudication while substantially outperforming ICD-based phenotyping approaches, particularly for disease severity classification. These findings demonstrate the potential of LLMs to transform unstructured clinical documentation into scalable, research-ready phenotypic data for large-scale glaucoma cohort development and EHR-based ophthalmic research.
Simons, G. J.; von Fersen, M.; Dahlberg, A.; Vartiainen, V.; Summanen, P.; Harju, M.
Show abstract
Background/Aims: Neovascular glaucoma (NVG) is a severe, secondary glaucoma. This study aimed to identify factors associated with vision, intraocular pressure (IOP), and ocular pain outcomes. Methods: The cohort included all patients diagnosed with NVG during 2008-2024 at Helsinki University Hospital, Finland. Linear mixed-effects models used pre-specified covariates, whereas machine learning was given the full longitudinal data with biomicroscopic findings as an exploratory approach. Results: 626 patients were analysed. Worse baseline vision and a closed angle were associated with worse follow-up vision. Treatments were associated with lower IOP and less pain rather than better vision. Age, sex and comorbidity were largely not associated with the outcomes. Glaucoma drainage devices showed the greatest initial IOP reduction (-10.2 mmHg, 95% confidence interval, CI -11.9 to -8.6 mmHg), followed by transscleral cyclophotocoagulation (TSCPC, -4.7 mmHg, 95% CI -5.8 to -3.7 mmHg) and peripheral retinal cryotherapy (-2.2 mmHg, 95% CI -3.1 to -1.4 mmHg). TSCPC and cryotherapy were also associated with reduced pain (odds ratio 0.51 and 0.46). Pan-retinal photocoagulation and anti-VEGF showed smaller IOP reductions, with a pain reduction for pan-retinal photocoagulation only. Both methods agreed, and machine learning added no novel clinical findings. Conclusions: Vision in this cohort was largely set by the state of the eye at diagnosis. IOP control and pain relief therefore remain realistic goals even when sight cannot be saved. Peripheral retinal cryotherapy stood out, linked to both lower IOP and less pain, seldom reported in NVG. These associations from a large, unselected cohort identify treatments worth comparing prospectively.
Ahmed, M. S.; Islam, M.; Albert, A.; Jahan, A. F.; Islam, N. L.; Shahnaz, T.; Mahdee, C. M.
Show abstract
Objective: To describe trends in cataract surgical volume, district-level surgical burden, and early postoperative visual outcomes among patients treated through a multi-district outreach programme in Bangladesh. Methods and Analysis: This retrospective study was conducted using data from patients undergoing cataract surgery through the outreach eye-camp programme of Bashundhara Eye Hospital and Research Institute across eight districts of Bangladesh, between 2016 and 2025 (excluding 2021 because of COVID-19). Annual surgical volume trend was assessed using Poisson regression. Postoperative visual outcome on day 1 was categorized as good, borderline, or poor per WHO criteria. Univariable and multivariable ordinal logistic regression identified predictors of worse outcome. Results: 1,929 cataract-surgery records were included. Surgical volume rose from 34 cases in 2016 to a peak of 677 in 2023 (IRR = 1.20; 95% CI: 1.180, 1.220; p-value< 0.001). Small incision cataract surgery (SICS) was used in 99.43% cases. On postoperative day 1, 70.09% of eyes had a good outcome, 19.44% borderline, and 10.47% poor. Increasing age was independently associated with worse outcome, with 2 to 3 times higher odds among patients over 70. District was independently associated with outcome, with Chapainawabganj and Kushtia having lower odds of worse outcome than Brahmanbaria. Sex was significant only in unadjusted analysis. Conclusion: Surgical volume rose substantially over time. About seven in ten eyes achieved a good outcome on day 1, with age and district as the main predictors of worse outcome. Limitations include a single early assessment, exclusion of incomplete records, no standardized refraction, and unmeasured predictors.
Liu, Z.; Fan Gaskin, J. C.; Ang, G. S.; Bigirimana, D.; Kong, G. Y. X.; Atik, A.; McGuinness, M. B.
Show abstract
Purpose The direct effect of iStent inject on intraocular pressure (IOP) in patients with glaucoma is difficult to quantify in pragmatic trials where rates of post-surgical IOP-lowering therapy differ between intervention groups. We aimed to quantify the causal effect of iStent inject on unmedicated IOP at 12- and 24-months post-surgery. Methods Adults with mild-to-moderate glaucoma were 1:1 randomised to receive cataract surgery with iStent inject or cataract surgery alone at an Australian hospital (2017-2020, NCT03106181). IOP-lowering medications were prescribed as per clinician discretion. An exploratory analysis was used to estimate the controlled direct effect of iStent inject on IOP, analogous to the effect expected if all participants had undergone medication washout prior to assessment. Results Ninety-five eyes from 80 people were included (67.4% male, mean age 73.0 years, mean baseline IOP 17.1 mmHg). IOP-lowering medication was required for 53% of eyes in each group at 12 months (n=76); at 24 months (n=86) it was required for 43% and 64% in the active and control groups, respectively. Mean IOP was similar between intervention groups at each outcome visit. The controlled direct effect favoured the iStent inject group at 12 months (-2.1-mmHg difference, 95% CI -4.0,-0.3) but was attenuated at 24 months (-0.5 mmHg-difference, 95% CI -2.6,1.6). Conclusion Although the iStent inject was estimated to have an effect on lowering unmedicated IOP at 12 months, this effect had largely disappeared by 24 months. Medication washout is recommended when safe and practical in future trials to estimate these direct effects with more certainty.
Shirzada, A.; Vlug, L.; Marinkovic, M.; Luyten, G. P. M.; Bleeker, J. C.; Vu, T. H. K.; Rasch, C. R. N.; Horeweg, N.; Pieterse, A.
Show abstract
Background: A subset of uveal melanomas can be treated using either enucleation or proton beam therapy (PBT), which offer similar oncological outcomes. The most appropriate treatment depends on a patient's preference. To allow patients to genuinely determine their preference, it is recommended to describe options as neutrally as possible. This study assesses to what extent ocular oncologists use and perceive non-neutral framing behaviour, and if it is related to patient satisfaction with decision-making. Methods: Consultations of ocular oncologists with patients newly-diagnosed with uveal melanoma were audio recorded, transcribed verbatim, and coded for ocular oncologists' explicit and implicit non-neutral framing behaviours. Explicit non-neutral framing was defined as: explicitly mentioning a preferred option at least once, without relating it to the patient. Implicit non-neutral framing was defined as: describing an option (un)favourably, without providing a medically substantive clarification alongside. Results: 110 patients provided consent for the audio recordings. Non-neutral framing was found in 84% (n=92/110) of consultations. We found explicit behaviour in 38% (42/110) and implicit behaviour in 76% (84/110, median=1, range, 0-4) of consultations. The most frequent implicit framing was presenting options by positively or negatively emphasizing one option. Non-neutral framing behaviours were not significantly related to patient satisfaction with decision-making. Conclusion: This study shows that in most consultations some non-neutral framing was present, which did not impact patients' satisfaction with decision-making. Nonetheless, ocular oncologists should be aware that how they describe options may influence preferences in ways that do not align with the patient's values.
Adan-Castro, E.; Nunez-Amaro, C. D.; Villareal, J.; H. Islas, I.; Hernandez-Quijano, A.; Rodriguez-Chagoya,, B. E.; Garcia-Roa, M.; Lopez-Star, E.; Garcia-Franco,, R.; Robles-Osorio,, M. L.; Martinez de la Escalera, G.; Clapp, C.
Show abstract
Background/Objective: Diabetic macular oedema (DMO) is a leading cause for visual impairment primarily managed with intravitreal anti-VEGF agents such as ranibizumab (RBZ). Levosulpiride (LSP), a prokinetic medication, was recently repositioned as a safe oral treatment for naive DMO. Here, we investigated the adjuvant effect of oral LSP in combination with intravitreal RBZ injections for treating persistent DMO. Subjects/Methods: Double-blinded, dual-centre, phase 2 trial in patients with centre-involving DMO randomly assigned to be orally treated with placebo (15 patients, 18 eyes) or LSP (18 patients, 19 eyes) along with 3 successive (4 weeks apart) RBZ intravitreal injections and a 24-week follow-up. Results: Baseline best-corrected visual acuity (BCVA) improved (p[≤]0.04) at week 12 in both RBZ+placebo and RBZ+LSP, but improvement was maintained (p=0.009) at week 24 only in RBZ+LSP. In agreement, longitudinal changes from baseline in BCVA from weeks 12 to 24 defined superior (p=0.02) visual gains measured by the Area Under the Curve (AUC) in RBZ+LSP vs. RBZ+placebo. The baseline value of mean central foveal thickness (CFT) decreased (p[≤]0.002) in both groups at week 12 and CFT reduction was significant (p=0.006) at week 24 only in RBZ+LSP. Also, longitudinal changes from baseline in CFT resulted in a higher AUC reduction (p[≤]0.04) at weeks 4 to12 in RBZ+LSP vs. RBZ+placebo. No significant adverse side effects were detected. Conclusions: Adjunctive LSP showed functional and anatomical benefits over the first-line therapy with RBZ. Adjuvant properties may involve the LSP-induced intraocular upregulation and downregulation of vasoinhibin and VEGF, respectively. Larger clinical trials are warranted.
Solages, N.; Scherer, R.; Samico, G. A.; Gutkind, N. E.; Kang, J.; Medeiros, F. A.; Swaminathan, S. S.
Show abstract
Purpose: To evaluate the efficacy of large language models (LLMs) in extracting medication-related information from glaucoma clinical notes in the electronic health record (EHR). Design: Cross-sectional. Subjects: 1,250 subjects in the Bascom Palmer Ophthalmic Repository. Methods: Extracted clinical notes from glaucoma-related encounters between 2014 and 2024 were labeled by two glaucoma specialists with a third serving as an adjudicator. Graders were asked to label current topical medications (CTM), proposed changes to topical medications ({Delta}TM), current oral medications (COM), and proposed changes to oral medications ({Delta}OM) in a structured fashion. The dataset was split into development (10%), validation (10%), and test (80%) sets stratified by clinician. Development and validation sets were used to engineer and refine prompts, and the held-out test set was used for model assessment. Five LLMs (Claude Opus 4.6, DeepSeek-V3.2, GPT 5.2, Grok 4.1, and Qwen3.6-35B-A3B) were accessed via Microsoft Azure AI Foundry within a HIPAA-compliant environment. Inter-grader agreement was assessed with Gwet AC1. LLM performance was initially assessed in a binary fashion with F1 scores, and the degree of text match among positive cases was evaluated using exact match accuracy and Jaccard Index (JI). Main Outcome Measures: F1 score, exact match accuracy, JI. Results: Gwet AC1 for intergrader agreement was 0.799, 0.888, 0.985, and 0.988 for CTM, {Delta}TM, COM, and {Delta}OM, respectively. F1 scores for CTM were 0.985, 0.971, 0.978, 0.968, and 0.970 for Claude, Deepseek, GPT, Grok, and Qwen, respectively; for {Delta}TM: 0.905, 0.826, 0.897, 0.842, 0.855, respectively; for COM: 0.923, 0.887, 0.899, 0.906, 0.894, respectively; for {Delta}OM: 0.958, 0.815, 0.937, 0.835, 0.940, respectively. Among positive cases, range of exact match accuracies for CTM (N=1354) was 0.730- 0.882 and range of JIs was 0.809-0.918. For {Delta}TM (N=404), exact match accuracy range was 0.619-0.780 and JI range was 0.668-0.827. For COM (N=47), exact match accuracy range was 0.766-0.872 and JI range was 0.765-0.870. For {Delta}OM (N=25), exact match accuracy range was 0.583-0.920 and JI range was 0.583-0.922. Conclusions: The GLLaucoMed pipeline demonstrated high performance in extracting and standardizing medication data from unstructured clinical notes, including both current medications and proposed changes. Claude and GPT exhibited the strongest performance.
Singh, A. M.; Yeh, T.-C.; DeBoer, C.; Al-Moujahed, A.; Lin, J. B.; Smith, S. J.; Sanislo, S.; Janjua, K. A.; Lin, T.-C.; Almeida, D. R. P.; Mruthyunjaya, P.; Mahajan, V. B.
Show abstract
Purpose: To evaluate the safety, procedural performance, sample recovery, and surgeon preference of an ophthalmic needle designed specifically for anterior chamber (AC) paracentesis. Methods: In this multicenter study, AC paracentesis was performed in clinic and operating-room settings using a 32-gauge x 4-mm needle with low dead space. The procedure was evaluated using a standardized physician survey. Prespecified outcomes included procedure-related adverse events (primary outcome), needle entry and handling, aspiration and sample recovery, comparative performance versus a 30-gauge needle, and physician preference for future use. Results: A total of 110 needle uses by eight surgeons were included. No ocular complications occurred, including lens or iris injury, hyphema, AC collapse, wound leak, hypotony, infection, or retinal complication, and no procedure required needle exchange or conversion to another device. Two technical events without ocular sequelae were noted, in which needle entry was partial thickness and did not reach the AC (1.8%; exact 95% CI, 0.2%-6.4%). Physicians rated needle entry, handling and sample recovery as good or excellent. Compared with a 30-gauge needle, the study needle was rated as at least comparable across all assessed domains. All surgeons rated it better or much better for intra-procedural safety and preferred it for future AC taps. Conclusions and Relevance: This short, 32-gauge low-dead-space ophthalmic needle demonstrated a favorable safety profile and was preferred over a 30-gauge needle by all surgeons. As aqueous humor liquid biopsy expands in clinical diagnostics and trials, an ophthalmic-specific needle design may help improve the consistency and safety of aqueous humor collection for molecular analysis and broader clinical use. Keywords: Anterior chamber paracentesis; Aqueous humor; Liquid biopsy; Low dead space; Ophthalmic needle
Jaurrieta Hinojos, J. N.; Palomares Ordonez, J. L.; Chacon Hinojos, J. F.; Folgueras Batres, M. A.
Show abstract
Abstract Background. Quantitative optical coherence tomography (OCT) measurements are essential for retinal disease monitoring, yet leading vendors store acquisition data in undocumented proprietary formats or encode measurements exclusively in private DICOM tags inaccessible to open systems. Methods. We present Transducin, an open-source Python library that reverse-engineers the undocumented Optopol Revo FC130 and Revo 60 .OPT binary format and extracts quantitative measurements from Zeiss Cirrus HDOCT private DICOM tags, generating TID 1500 Structured Reports with SNOMEDCT coded findings for both platforms. A novel finding, that OCTPARAMS tag 23 encodes ocular laterality through the arithmetic sign of the foveal horizontal position, enables geometry based laterality inference requiring no operator data entry, validated across 18 files from two device models and four software versions with 100% accuracy. Results. The primary corpus of 452 Optopol .OPT files (73 patients, 7 acquisition types) was parsed with 100% success. Cross-version compatibility was confirmed across SOCT versions 11.5.0 through 21.5.0, spanning approximately eight years of software development. The Zeiss Cirrus pipeline generated TID 1500 SRs for all 41 applicable studies (100%), yielding CMT 203to 630um and RNFL 53 to123 um across a clinically representative range. Conclusions. Transducin provides the first publicly documented specification of the Optopol .OPT format and the first open-source multivendor pipeline generating SNOMEDCT coded DICOM Structured Reports from both Optopol Revo and Zeiss Cirrus devices, closing a gap explicitly confirmed by both manufacturers' own documentation. The code is available at https://github.com/oftalmos-org/transducin (Apache License 2.0).
Aurilia, A.; Martin, N.-L.; Simon-Martinez, C.; Antoniou, M.-P.; Bouthour, W.; Bavelier, D.; Backus, B. T.; Dornbos, B.; Blaha, J. J.; Kropp, M.; Muller, H.; Murray, M. M.; Thumann, G.; Steffen, H.; Matusz, P. J.
Show abstract
Objectives: Amblyopia is a pediatric visual disorder traditionally treated by patching the fellow eye, though many patients retain residual amblyopia post-treatment. Increasing evidence suggests that visual plasticity allows treat-ment beyond the classical therapeutic window. AMBER evaluated the efficacy of binocular serious games in virtual reality (VR) in residual amblyopia. Methods and Analysis: The monocentric, prospective, randomized, crossover trial (reported as case series) includ-ed 14 anisometropic, strabismic, or mixed residual amblyopia patients (6-35 years; 5 children, 9 adults). Participants underwent two 2-month intervention phases: optical correction (standard care) and standard care plus VR games (2.5 h/week), each with a 2-month follow-up. Best-corrected visual acuity (BCVA), stereoacuity, and reading speed were assessed (5 timepoints) using the Sloan and Landolt charts, the Titmus, TNO, Lang II, Asteroid, and Mnread tests. Compliance and adverse events (AE) were recorded. Results: VR training improved BCVA in 10 amblyopic eyes (Landolt and Sloan), with more pronounced effects in anisometropic patients. Six patients showed improved stereoacuity (Titmus; 4x mixed, 1x anisometropic, 1x stra-bismic amblyopia), persistent only in children (1x strabismic, 1x mixed amblyopia). Four improvements were ob-served with TNO (1x), Lang II (1x), Asteroid (0x), and MNread (1x). Despite positive trends, when comparing re-sults of individual patients, between both eyes, and with standard treatment, consistency of improvements cannot be conclusively demonstrated. One non-severe AE (dizziness) was reported. Conclusions: Following individual cases, VR training improved BCVA and stereoacuity, particularly in children and patients with high compliance. However, considering the cohort as a whole, consistency of effects has to be confirmed in larger groups. Thus, the methodologically sophisticated AMBER study revealed differences in VR treatment efficacy between amblyopia types, children/adults, endpoints and tests, offering precious data for the design of meaningful future studies. It shows that neurovisual plasticity gauged by VR-games offers safe, engaging treatment options for residual amblyopia.
Bakaraju, R. C.; Bandela, P. K.; Sha, J.; Tilia, D.
Show abstract
Clinical relevance: Validated virtual control arms may provide population-level estimates of treatment effect and reduce reliance on untreated control allocations in myopia trials. Background: Untreated single-vision control arms in paediatric myopia efficacy trials are increasingly difficult to justify and retain. Several published models predict untreated childhood axial elongation by region or ethnicity. Here they are implemented unchanged in an open-source tool and validated against an untreated multi-ethnic cohort. Methods: Five published models predicted untreated elongation from baseline age, cycloplegic spherical equivalent, sex, and ethnicity, anchored at baseline axial length (AL) and evaluated at actual follow-up. Predictions were compared with 242 untreated myopic children (Chinese, Vietnamese, Indian) with AL measured at approximately 6 and 12 months, assessing bias, root-mean-square error, and prediction-interval coverage against pre-specified thresholds (bias <0.03 mm; coverage greater than or equal to 0.90). Results: The regional generalised estimating equation (GEE) and meta-regression models reproduced mean East Asian elongation without meaningful bias at 6 months (GEE bias -0.013 mm; equivalence to plus-or-minus 0.03 mm, p = 0.014) and at 12 months (-0.004 mm), although equivalence was not established at 12 months in an underpowered subgroup (n = 71, all Vietnamese; p = 0.068). Older age-only models under-predicted by 0.07 to 0.12 mm. Published individual prediction intervals were too narrow (coverage 0.77): the means were accurate, the individual uncertainty was not. Indian elongation fell between strata and was matched by no existing model. Conclusions: The models reproduce mean untreated East Asian elongation at 6 months, conditional on cohort independence; South Asian children remain unserved by any existing stratum. The tool is a group-level instrument, not an individual predictor, and a transparent unification of published models in open-source code. Its value for estimating treatment effect awaits back-testing against a trial with a known untreated arm, ideally over 24 to 36 months.
Jaurrieta Hinojos, J. N.; Gonzalez Saldivar, G.; Hernandez Vazquez, A. Y.; Saucedo Castillo, A.; Babayan Sosa, A.; Ramirez Estudillo, J. A.
Show abstract
Purpose: To assess the feasibility of quantitative fundus autofluorescence (FAF) measurement in early age-related macular degeneration (AMD) using the freely available ImageJ software, to characterize signal intensity across FAF patterns, and to evaluate interobserver reproducibility in pattern classification. Methods: Single-center, non-blinded, retrospective, consecutive-case analytical study. FAF images acquired with Spectralis OCT+HRA (Heidelberg Engineering) from patients with early dry AMD seen at a tertiary referral center between January 2010 and September 2016 were analyzed. A standardized 300x300-pixel region of interest (ROI) centered on the fovea was evaluated in ImageJ v2.0.0-rc54/1.51h (Fiji distribution). Mean, minimum, and maximum autofluorescence (AF) pixel intensity were recorded. Each image was independently classified according to the Bindewald classification system by two graders; a third senior grader adjudicated discordances. Cohen's kappa (k) was used to assess interobserver agreement. Results: Of 423 patients with available FAF studies, 107 had dry AMD; 45 met quality and diagnostic criteria for early AMD and were included in the quantitative analysis. Mean age was 73.47 +/- 8.1 years; 62.2% were female. Mean FAF intensity was 120.26 (range 74.76-160.79); mean minimum was 32.07 (range 3-63) and mean maximum was 205.80 (range 125-255). Seven of eight Bindewald patterns were identified; the stippled pattern was absent. The most frequent pattern was minimal changes (31.1%), followed by increased focal (24.4%) and patchy (15.6%). Reticular pattern showed the highest mean AF (143.8), while lace pattern showed the lowest (88.4). Interobserver agreement for Bindewald pattern classification was almost perfect (k = 0.969; 95% CI, 0.908-1.000; p < 0.001). Agreement for lesion extent was moderate (k = 0.531) and for foveal involvement was substantial (k = 0.622). Conclusions: Quantitative FAF evaluation of early AMD using ImageJ is feasible and reproducible. ImageJ represents a cost-free alternative for multimodal retinal image analysis, with potential for automated screening applications in resource-limited settings. Keywords: age-related macular degeneration; fundus autofluorescence; ImageJ; quantitative autofluorescence; image analysis; Bindewald classification; interobserver agreement
Ragni, F.; Bresolin, P.; Cocu, M.; Bovo, S.; Malfatti, G.; Cagol, D.; Inchiostro, S.; Romanelli, F.; Moroni, M.; Jurman, G.
Show abstract
Diabetic retinopathy (DR) screening commonly relies on fixed follow-up intervals, although progression risk differs across patients. We developed a preliminary image-clinical framework to support personalized follow-up recommendations from retinal fundus images and systemic risk factors. Public DR datasets were harmonized into a binary task distinguishing absence of DR from DR of any grade. An ImageNet-pretrained ResNet50 and a foundation model-based feature extraction pipeline were compared. The selected image model was integrated with literature-derived severe retinopathy progression curves and clinical risk modifiers to estimate personalized cumulative risk and assign follow-up intervals using a predefined acceptable risk threshold. The ResNet50-based model was selected for subsequent analyses. In the target cohort, the integrated model assigned 90.0\% of patients to follow-up within 12 months. Clinical adjustment substantially modified image-only recommendations, generally shifting patients toward shorter intervals. These preliminary findings support the feasibility of combining image-derived estimates of baseline DR status with clinical modifiers to inform personalized screening intervals. Larger longitudinal studies are needed to validate calibration, clinical utility, and safety before real-world implementation.
Calder, D.; Johnson, T.; Carpenter, C.; Tan, C.; Omotowa, O.; Wu, C.; Singer, P.; Hu, K.; Stagg, B.
Show abstract
Abstract Importance: Access to eye care is increasingly constrained by declining ophthalmologist workforce density, particularly in rural areas, while optometrist workforce is projected to exceed demand. Numerous state legislatures have expanded optometrist scope of practice (SOP), but the workforce effects of these policies have not been evaluated.. Objective: To evaluate changes in optometrist and ophthalmologist workforce density following state-level expansion of optometrist scope of practice. Design: This retrospective ecological study analyzed workforce density across 50 states and the District of Columbia. Between 2008 - 2019, nine states expanded SOPs for optometrists. We used interrupted time-series regression to estimate associations between these policy changes and provider density from 2010-2021, adjusting for covariates. Setting: Population-based analysis of all 50 US states and the District of Columbia, 2010 - 2021. Participants: State-level workforce data were derived from Bureau of Labor Statistics and US Census data (optometrists) and the American Medical Association Physician Masterfile (ophthalmologists). Socioeconomic covariates were obtained from the American Community Survey. Exposure: State-level legislative expansion of optometrist SOPs to allow injections beyond anti-anaphylaxis treatment, lesion removal, and/or laser procedures. SOP changes were identified through systematic review of legislative and regulatory records. Main Outcomes and Measures: Annual change in optometrist and ophthalmologist workforce density (providers per 100,000 population) following SOP expansion, adjusted for age, income, rates of vision difficulty, diabetes, poverty, and uninsured status. Results: Nine states met inclusion criteria for SOP expansion between 2008 and 2019. Analysis of national data showed that among all 50 US states and the District of Columbia, SOP expansion was associated with a non-significant change in optometrist density (-0.65 per 100,000; 95% CI, -1.87 to 0.57) and ophthalmologist density (+0.09; 95% CI, -0.10 to 0.28). Results were consistent in the 9-state subgroup (optometrists: -0.60, 95% CI -2.90 to 1.70; ophthalmologists: +0.02, 95% CI -0.13 to 0.18). Conclusions and Relevance: Our population-level data suggest that, in the US, state legislation expanding optometrist scope of practice has not been associated with increased workforce density. As numerous state legislatures continue to consider such policies, these findings can inform efforts to balance access to eye care with patient safety.
zhou, k.; chen, y.; YILDIZ, E.; Shi, M.; Dai, D.; Chen, G.; Zheng, J.; Wang, H.; Zhan, F.; Saini, C.; Shen, L. Q.; Guo, Y.; Liang, P. P.; Wang, M.
Show abstract
Glaucoma is a leading cause of irreversible blindness worldwide. Ophthalmologists diagnose glaucoma through a structured reasoning process by sequentially evaluating optic nerve head characteristics before reaching a final diagnosis, whereas existing AI systems typically perform direct image classification without providing clinically meaningful reasoning. We present the first clinically annotated fundus reasoning dataset, comprising 1,077 fundus photographs paired with expert-authored six-step diagnostic reports. Building on this dataset, we develop a reasoning-driven vision-language framework that explicitly models the ophthalmologist's diagnostic workflow by generating structured clinical reasoning prior to diagnosis. The generated reports are clinically validated, achieving the best performance across all evaluated clinical findings, including a cup-to-disc ratio mean absolute error of 0.070, an ISNT Kendall distance of 1.73, and the highest semantic agreement with expert reports (BERTScore-F1 = 0.874). The resulting framework also improves glaucoma diagnosis, achieving a balanced accuracy of $94.7\%$ and precision of $94.8\%$, demonstrating that explicitly modeling expert clinical reasoning simultaneously improves interpretability and diagnostic performance. Code and data are available at \url{https://glaucoma-cot.github.io/}.
Murphy, T. I.; Armitage, J. A.
Show abstract
Purpose: To investigate how artificial intelligence (AI) systems detect referrable diabetic retinopathy (DR) from retinal photographs by analysing heatmap patterns and determining their overlap with DR features. Methods: Fifty-four AI systems were developed using 27 backbone architectures, with each implemented as both binary-referable and multi-class grading models based on the International Clinical Diabetic Retinopathy (ICDR) grading scale. Models were trained on images from DDR, BRSET and Kaggle datasets. After training, each model analysed 749 images with DR feature annotations, with Grad-CAM heatmaps generated and compared to pixel-level annotations of microaneurysms, haemorrhages, exudates, cotton wool spots, venous beading, intraretinal microvascular abnormalities and neovascularisation. Results: All models achieved acceptable predictive performance (AUROC >0.8 for most architectures). Heatmap analysis revealed consistent attention to the macular region with relative neglect of the optic disc. Exudates and cotton wool spots were highlighted most frequently by the heatmaps, with venous beading and neovascularisation at the disc showing poor overall coverage for binary referable classifiers. Models grading per the ICDR scale demonstrated high coverage for all features. Substantial variability was observed between architectures, suggesting different feature detection capabilities. Interestingly, the heatmap analysis indicated that the models were using different logic to the ICDR grading scale definitions. Conclusion: AI models do not uniformly rely on all DR features when detecting referable DR, limiting their predictive performance in unusual presentations. Heatmap aggregation analysis provides a scalable method for analysing model behaviour, allowing strengths and weaknesses to be identified. These findings may help improve clinician's trust and acceptance of AI.
Ahmed, M. S.; Islam, M.; Albert, A.; Jahan, A. F.; Islam, N. L.; Shahnaz, T.; Mahdee, C. M.
Show abstract
Introduction: Community outreach eye camps extend eye care to underserved rural populations across South Asia, but published data mostly describe surgical yield rather than the full range of presenting conditions. This study described the pattern of ocular diagnoses among patients attending community eye camps in Bangladesh and examined how these patterns varied by year, age, and sex. Methodology: This was a retrospective, registry-based study of patients screened through the outreach eye-camp programme of Bashundhara Eye Hospital and Research Institute, Bangladesh, between 2020 and 2024 (2021 excluded due to COVID-19). Age, sex, and provisional diagnosis were recorded for every patient and grouped into five categories: cataract and lens disorders, refractive errors and presbyopia, ocular surface inflammatory disorders, lacrimal system disorders, and miscellaneous disorders. Multivariable logistic regression models estimated the adjusted odds and predicted prevalence of each diagnostic category by calendar year (adjusted for age group and sex) and by age group and sex (adjusted for calendar year). Results: Among 7,267 patients, refractive errors and presbyopia were the most common category (51.31%), followed by cataract and lens disorders (23.54%), miscellaneous disorders (13.86%), ocular surface inflammatory disorders (8.59%), and lacrimal disorders (2.70%). Ocular surface disorders declined significantly over time (adjusted OR: 0.80/year; p-value<0.001). Cataract showed a borderline decline (adjusted OR: 0.94; p-value=0.065). Cataract prevalence rose steeply with age, from 3-6% below 18 years to 49-55% at 60 and older, while refractive errors/presbyopia peaked at 40-59 years (64-68%) before declining thereafter (30-37%). Ocular surface disorders were most frequent among children. Lacrimal disorders remained uncommon (2-5%), with no significant interaction. Conclusion: Refractive errors and presbyopia were the leading reason for presentation, with distinct age patterns across diagnostic categories. Outreach eye camps should be resourced for both comprehensive refraction, primary eye care and cataract-surgical referral, reflecting actual community-level diagnostic needs.
Nagalamadaka, P.; Ross, C. J.; Gilbert, J. B.; Stillman, H.; Ghauri, S. Y.; Dutton, S. M.; Kearney, W.; Li, J. H.; Leong, A.; Singh, R. P.; Krzystolik, M. G.
Show abstract
Purpose: To evaluate whether initiation of GLP-1 receptor agonists (GLP-1RAs) is associated with anti-VEGF treatment burden in type 2 diabetes patients with diabetic macular edema (DME) in the IRIS(R) Registry (Intelligent Research in Sight). Methods: Incident GLP-1RA initiators were matched 1:1 with controls via Mahalanobis distance matching (9,896 pairs; N=19,792) on sociodemographics, DME risk factors, and factors influencing GLP-1RA prescription including hypertension, obesity, chronic kidney disease. A longitudinal mixed-effects event-study model evaluated monthly anti-VEGF injection frequency over a 36-month window (12 months before through 24 months after initiation), adjusting for DME duration. Visual acuity (VA) and central subfield thickness (CST) were secondary outcomes. Results: Following GLP-1RA initiation, anti-VEGF injection trajectories did not significantly differ between the matched GLP-1RA and control cohorts (interaction coefficients -0.18 to 1.59, P>0.05). Likewise, no differences in VA were observed between cohorts (-0.05 to 0.04 logMAR, P>0.05) or CST (-14.12 to 33.58 {micro}m, P>0.05). Conclusion: In these matched cohorts, GLP-1RA initiation was not associated with the trajectory of anti-VEGF use or changes in VA or CST. Precis We used the American Academy of Ophthalmology IRIS(R) Registry (Intelligent Research in Sight) to identify patients with DME. In 19,792 matched patients, there was no significant reduction in injection frequency post GLP1-RA initiation and no significant change in VA or CST.
Sahoo, N. K.; Doshi, U.; Gregori, G.; Flores-Pena, D.; Lupidi, M.; Vupparaboina, K. K.; Chhablani, J.
Show abstract
Purpose: To validate an automated pipeline to detect and quantify focal retinal and choroidal pulsation areas that are synchronous with the cardiac cycle in video indocyanine green angiography (ICGA). Design: Retrospective, observational, hypothesis-generating validation study Subjects, Participants: Consecutive patients with a diagnosis of central serous chorioretinopathy (CSCR) in one or both eyes. Methods: Videos were acquired on Heidelberg HRA+OCT. The pipeline consisted of three steps: signal extraction, foci detection, and quantification. After registration of the constituent frames, each pixel's intensity signal was analyzed at the presumed cardiac frequency (tested from a sample of three detectable frequencies). A synchrony score combining local phase coherence with oscillation amplitude was then derived and computed using a standard deviation ({sigma}) above each video's background oscillation value. Two masked graders marked the retinal and choroidal pulsation areas twice. We compared detection of the pulsation areas against grader consensus using a receiver operating characteristic curve (using multiple grid sizes to divide the scan area) and, separately, using a signal-based area-reduction method to obtain an optimum {sigma} value. Main Outcome Measures: Agreement between the automated algorithm and human graders in detection of pulsation foci, and the optimum threshold multiplier ({sigma}). Results: We studied 20 ICGA videos from 20 eyes. At the 16-pixel grid size, the pipeline achieved a mean area under the curve (AUC) of 0.914, sensitivity of 0.86, and specificity of 0.80. Grader agreement improved with larger grid size, reaching substantial-to-strong levels for choroidal annotations. The two independent validation methods demonstrated similar {sigma} values that differed by 0.62{sigma}, supporting {sigma}=4.0 as the optimum value. Conclusions: We report the first automated method to quantify retinal and choroidal vascular pulsation on video ICGA. It measures pixels that oscillate over time with the presumed cardiac cycle and works reliably at the spatial scale (grid level) where experts agree. Pulsatile hemodynamics may add a new vascular biomarker for glaucoma, diabetes, hypertension, and pachychoroid diseases.
Moin, M.; Younas, R.; Maqbool, S.; Awan, Z. H.; Bilal, M.; Katibeh, M.; Watts, E.; Latorre-Arteaga, S.; Freels, P. E.; Steinberg, A.; Bastawrous, A.
Show abstract
Background: An estimated 826 million people have avoidable near vision impairment (NVI) due to lack of access to near vision glasses for presbyopia. Glasses use, uptake and replacement is essential to sustained NVI correction. Here we explore the uptake of second and subsequent pairs of near vision glasses and the willingness to pay. Methods: In this cross-sectional survey, 276 individuals who received near vision glasses through eye health screening programs in Punjab province, Pakistan from 2020 to 2022, were surveyed between February and July 2025, i.e. 3-5 years later. Results: 92.0% (n=254) of respondents were still using near vision glasses: a purchased replacement (41.3%, n=114), a free replacement (30.8%, n=85), or the original pair (19.9%, n=55). Male gender, higher age, and personal income were strongly associated with purchasing additional pairs, while women and those economically dependent on others relied more on free provision (all p values <0.05). Accessibility, indicated by shorter travel times and awareness of supply sources, also played a significant role in sustained replacement. Over 98% of participants expected to obtain (purchase or receive) a new pair of glasses in the future, and over 94% of participants expected to purchase their next pair. The mean willingness-to-pay (WTP) was 353 PKR (95% CI: 327-379) // 1.25 USD (95% CI: 1.16-1.34), while the average reported price paid was 416.67 PKR // 1.47 USD (95% CI: 351.39-481.94)). Conclusion: Sustained use of near vision glasses remained high 3-5 years after initial provision. Many recipients had purchased replacements, and willingness and ability to pay were high among participants, suggesting potential for sustained demand beyond initial subsidized distribution.